Computer Science Notes: Mathematical Underpinnings of LLMs (Tokenization & Vectors)
Textual Discretization (Tokenization)
To comprehend Large Language Models (LLMs) beyond superficial perceptions of artificial sentience, one must examine the rigorous mathematical transformations executing within high-dimensional vector spaces.
- Token Decomposition: Words and subwords are mapped to discrete numerical token IDs. A single word may constitute one token or split into smaller subword chunks.
- Context Window: The rigid, upper-bounded memory buffer defining the maximum sequence of past tokens the attention mechanism can retain to compute the subsequent probability matrix.
Vector Embeddings & Semantic Topography
- High-Dimensional Vector Representation: Every token ID is transformed into a dense floating-point vector tensor (frequently spanning 1536, 4096, or more dimensions).
- Semantic Proximity (Euclidean & Cosine Similarity): Conceptually synonymous or related entities cluster geometrically adjacent to one another within this multi-dimensional space.
===================================================================================
HIGH-DIMENSIONAL EMBEDDING VECTOR SPACE
===================================================================================
[ Cat ] (0.25, 0.89, 0.12) <ββ Close Cosine Distance ββ> [ Dog ] (0.28, 0.85, 0.15)
<ββ Massive Distance ββββ> [ Car ] (0.91, 0.05, 0.88)
===================================================================================
The Self-Attention Mechanism
Allows the model to dynamically compute query, key, and value (Q, K, V) correlation matrices between every token and all surrounding tokens, enabling flawless resolution of complex syntactic dependencies and ambiguous pronouns.